book summary
Evaluating book summaries from internal knowledge in Large Language Models: a cross-model and semantic consistency approach
We study the ability of large language models (LLMs) to generate comprehensive and accurate book summaries solely from their internal knowledge, without recourse to the original text. Employing a diverse set of books and multiple LLM architectures, we examine whether these models can synthesize meaningful narratives that align with established human interpretations. Evaluation is performed with a LLM-as-a-judge paradigm: each AI-generated summary is compared against a high-quality, human-written summary via a cross-model assessment, where all participating LLMs evaluate not only their own outputs but also those produced by others. This methodology enables the identification of potential biases, such as the proclivity for models to favor their own summarization style over others. In addition, alignment between the human-crafted and LLM-generated summaries is quantified using ROUGE and BERTScore metrics, assessing the depth of grammatical and semantic correspondence. The results reveal nuanced variations in content representation and stylistic preferences among the models, highlighting both strengths and limitations inherent in relying on internal knowledge for summarization tasks. These findings contribute to a deeper understanding of LLM internal encodings of factual information and the dynamics of cross-model evaluation, with implications for the development of more robust natural language generative systems.
CLIPPER: Compression enables long-context synthetic data generation
Pham, Chau Minh, Chang, Yapei, Iyyer, Mohit
LLM developers are increasingly reliant on synthetic data, but generating high-quality data for complex long-context reasoning tasks remains challenging. We introduce CLIPPER, a compression-based approach for generating synthetic data tailored to narrative claim verification - a task that requires reasoning over a book to verify a given claim. Instead of generating claims directly from the raw text of the book, which results in artifact-riddled claims, CLIPPER first compresses the book into chapter outlines and book summaries and then uses these intermediate representations to generate complex claims and corresponding chain-of-thoughts. Compared to naive approaches, CLIPPER produces claims that are more valid, grounded, and complex. Using CLIPPER, we construct a dataset of 19K synthetic book claims paired with their source texts and chain-of-thought reasoning, and use it to fine-tune three open-weight models. Our best model achieves breakthrough results on narrative claim verification (from 28% to 76% accuracy on our test set) and sets a new state-of-the-art for sub-10B models on the NoCha leaderboard. Further analysis shows that our models generate more detailed and grounded chain-of-thought reasoning while also improving performance on other narrative understanding tasks (e.g., NarrativeQA).
Book summaries made by Artificial Intelligence - How smart Technology changing lives
An Artificial Intelligence is capable of achieving many things, but when there is text interpretation or creation from scratch, things get more difficult. Still, it is not a difficult task, and with the example that I present today I show it. It is possible to create entire book summaries using a system from OpenAI, founded by Elon Musk, an intelligence capable of finding events in a book and making truly impressive summaries. They demonstrated this with "Alice's Adventures in Wonderland", a book of more than 26,000 words that was reduced to 6,000, although they also did so with "Romeo and Juliet" and "Pride and Prejudice." Still, it is only a first step to something much bigger, since for now humans have to analyze the results and make corrections so that the model continues to learn.
Book Summary: Life 3.0: Being Human in the Age of A.I. by Max Tegmark
The authors purpose for this book is to acknowledge this uncertainty and prompts us to collectively make some choices now. The book begins with a prelude which is a story of a possible near-future. I found it so fascinating that I have copied it out in full (it's just over 6000 words so its a 20 minute read). Have you seen Netflix's "Black Mirror"? The prelude reminds me of an episode of that show which explores possible technological futures which are frighteningly plausible.